Papers with vision-and-language models
Leverage Points in Modality Shifts: Comparing Language-only and Multimodal Word Representations (2023.starsem-1)
Copied to clipboard
| Challenge: | a recent study of the effect of visual grounding on language representations has given a new life to the debate around extractability and quality of semantic information in representations trained solely on textual input. |
| Approach: | They compare word embeddings from vision-and-language models to text-only models . they identify meaning properties and relations that characterize words whose embeddements are most affected by visual grounding . |
| Outcome: | The proposed model differs from text-only models on semantic representations of language . the study is the first large-scale study of the effect of visual grounding on language representations . |
Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work has focused on understanding the robustness of vision-and-language models to imperceptible variations on benchmark tasks. |
| Approach: | They develop a model that generates additional dilution text that maintains relevance and topical coherence with the image and existing text, and when added to the original text, leads to misclassification of the multimodal input. |
| Outcome: | The proposed model outperforms fusion-based classifiers on Crisis Humanitarianism and Sentiment Detection tasks by 23.3% and 22.5% in presence of dilutions generated by the model. |
Combining Tradition with Modernness: Exploring Event Representations in Vision-and-Language Models for Visual Goal-Step Inference (2023.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for representing procedural knowledge are limited to capturing the most crucial information, namely actions and the participants, to learn stereotypical event sequences. |
| Approach: | They propose a task that uses images to identify steps towards achieving a goal in the multimodal domain. |
| Outcome: | The proposed task uses images that represent steps towards achieving a textually expressed goal in the multimodal domain. |
Visual Spatial Reasoning (2023.tacl-1)
Copied to clipboard
| Challenge: | Existing benchmarks for testing vision-language models (VLMs) are not ideal as they conflate multiple sources of error and do not allow controlled analysis on specific linguistic or cognitive properties. |
| Approach: | They present a dataset containing more than 10k natural text-image pairs with 66 types of spatial relations in English (e.g., under, in front of, facing). |
| Outcome: | The proposed model fails to capture relational information in a visual question answering task and referring expression comprehension tasks. |
Semantically Distributed Robust Optimization for Vision-and-Language Inference (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to integrate linguistic knowledge into training pipelines are under-explored. |
| Approach: | They propose a model-agnostic method that leverages linguistic transformations to infer a set of linguistic phenomena. |
| Outcome: | The proposed method improves on benchmark datasets with images and video and is generalizable to other V&L tasks. |
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing vision-and-language models perform better on multimodal tasks, but there is little understanding of how multimodal learning can help visual representations. |
| Approach: | They conduct a probing analysis of visual representations in existing vision-and-language models and vision-only models by probing on a broad range of tasks. |
| Outcome: | The proposed model improves vision-and-language models on label and attribute prediction tasks while vision-only models are stronger on dense prediction tasks. |
Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation (2021.acl-long)
Copied to clipboard
| Challenge: | Edited media frames are structured annotations with respect to intents, emotional reactions, attacks on individuals, and the implications of disinformation. |
| Approach: | They propose a new formalism to understand visual media manipulation as structured annotations with respect to intents, emotional reactions, attacks on individuals, and the implications of disinformation. |
| Outcome: | The proposed model obtains promising results on a dataset with 56k question-answer pairs written in rich natural language. |
Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Generally, commonsense knowledge is correlated with culture and geographic locations and is only shared locally. |
| Approach: | They construct a Geo-Diverse Visual Commonsense Reasoning dataset to test vision-and-language models’ ability to understand cultural and geo-location-specific commonsense. |
| Outcome: | The proposed models perform better in non-Western regions including East Asia, South Asia, and Africa than in the Western regions. |
Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Pretrained machine learning models perpetuate and even amplify existing biases in data . this can result in unfair outcomes that ultimately impact user experience . |
| Approach: | They quantify bias amplification in pretraining and after fine-tuning on vision-and-language models. |
| Outcome: | The results show that pretrained models can perpetuate and even amplify biases in data without compromising performance. |
VAQUUM: Are Vague Quantifiers Grounded in Visual Data? (2025.findings-acl)
Copied to clipboard
| Challenge: | a dataset containing 20,300 human ratings on quantified statements is used to evaluate the appropriateness of vague quantifiers in visual contexts. |
| Approach: | They use a visual-language-models-based dataset to evaluate the appropriateness of vague quantifiers. |
| Outcome: | The proposed model is based on a visual-visual-language-model-based dataset . it shows that the model is compatible with humans when producing or judging vague quantifiers . |
Transparent and Coherent Procedural Mistake Detection (2025.emnlp-main)
Copied to clipboard
| Challenge: | Procedural mistake detection (PMD) is a problem of classifying whether a human user has successfully executed a task. |
| Approach: | They extend PMD to require generating visual self-dialog rationales to inform decisions . they leverage a natural language inference model to formulate two automated metrics for coherence of generated rationale. |
| Outcome: | The proposed model improves on a reframed task with a natural language inference model and a multi-faceted metrics visualization of common outcomes. |